Papers with evaluation protocols

41 papers
Knowledge Control for Responsible Generative AI: Bridging Academia, Industry, and Society (2026.acl-tutorials)

Copied to clipboard

Challenge: This tutorial introduces the foundations of post-training knowledge control and showcases recent frontier methods.
Approach: This tutorial introduces the foundations of post-training knowledge control and showcases recent frontier methods.
Outcome: This tutorial introduces the foundations of post-training knowledge control and showcases recent frontier methods . key motivations and failure modes, harmful generation and stereotype reinforcement, are addressed . core methods such as machine unlearning, knowledge editing, and inference-time interventions are also included .
Distribution Shifts Are Bottlenecks: Extensive Evaluation for Grounding Language Models to Knowledge Bases (2024.eacl-srw)

Copied to clipboard

Challenge: Existing benchmarks fail to reflect robustness challenges and fairly evaluate models.
Approach: They propose to ground language models to knowledge bases to investigate distribution shifts in language and linguistic aspects of distribution shift.
Outcome: The proposed method fails to evaluate language models in large and small datasets . the proposed model fails to cope with unseen schemas and language variations .
Establishing Trustworthiness: Rethinking Tasks and Model Evaluation (2023.emnlp-main)

Copied to clipboard

Challenge: Language understanding is a multi-faceted cognitive capability, which the Natural Language Processing community has striven to model computationally for decades.
Approach: They propose to rethink what constitutes tasks and model evaluation in NLP and pursue a more holistic view on language, placing trustworthiness at the center.
Outcome: The proposed models are based on generative models and are being deployed in more real-world scenarios, including previously unforeseen zero-shot setups.
Can Pretrained Language Models Derive Correct Semantics from Corrupt Subwords under Noise? (2023.starsem-1)

Copied to clipboard

Challenge: Existing studies have shown that Pretrained Language Models (PLMs) perform poorly under noise due to subword segmentation.
Approach: They propose a framework for subword segmentation that provides a systematic categorization of segmentation corruption under noise and evaluation protocols by generating contrastive datasets with canonical-noisy word pairs.
Outcome: The proposed framework provides a systematic categorization of segmentation corruption under noise and evaluation protocols by generating contrastive datasets with canonical-noisy word pairs.
Interpreting Predictive Probabilities: Model Confidence or Human Label Variation? (2024.eacl-short)

Copied to clipboard

Challenge: In modern NLP, neural networks are the de-facto standard to predict complex probability measures from available context.
Approach: They propose to use a single predictive distribution to evaluate models with disentangled representations of uncertainty about predictions and uncertainty about human labels.
Outcome: The proposed models are crucial for trustworthy and fair NLP systems, but exploiting a single distribution is limiting.
Fake News Detection Strategies under Dataset Bias: Using Large-scale Coarse-grained Labels (2026.eacl-srw)

Copied to clipboard

Challenge: Existing datasets differ substantially in content distributions and annotation policies, complicating fair evaluation and generalization assessment.
Approach: They quantitatively analyze dataset bias across multiple public fake news datasets with different annotation granularities, including article-level and publisher-level labels.
Outcome: The proposed approach improves detection performance under in-dataset and cross-data set evaluation settings.
Garbage In, Reasoning Out? Why Benchmark Scores are Unreliable and What to Do About It (2026.findings-eacl)

Copied to clipboard

Challenge: Using social reasoning benchmarks, we uncover pervasive flaws in both benchmark items and evaluation methodology.
Approach: They audit three widely used social reasoning benchmarks and identify flaws in their design and evaluation methodology.
Outcome: The results challenge the validity of current benchmark-based claims about social reasoning in large language models.
Quantifying Aleatoric Uncertainty of In-Context Learning for Robust Measure of LLM Prediction Confidence (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for standard generation tasks fail to capture the unique dynamics of ICL.
Approach: They propose a concept of self-function vectors that leverage Bayesian views and the mechanistic interpretability of ICL to model latent concept learned during in-context prompting.
Outcome: The proposed framework can be used for trustworthy-related applications, such as hallucination detection.
Thesis Proposal: A Normalization-First Framework for Sound, Complete, and Utility-Ready Open Information Extraction (2026.acl-srw)

Copied to clipboard

Challenge: Existing approaches to extract relational tuples from text are incomplete and ambiguous . Existing methods rely on predefined schemas to produce t-uples .
Approach: They propose a normalization-first framework that reframes OIE as a structured semantic transformation pipeline . they formalize soundness, completeness, and usefulness as approximate yet verifiable guarantees over extraction quality .
Outcome: The proposed framework aims to make OIE usable for downstream reasoning and machine interpretability.
EVI: Multilingual Spoken Dialogue Tasks and Dataset for Knowledge-Based Enrolment, Verification, and Identification (2022.findings-naacl)

Copied to clipboard

Challenge: Knowledge-based authentication is crucial for task-oriented spoken dialogue systems that offer personalised and privacy-focused services . e-learning systems should be able to enrol, identify, and verify new and recurring users based on their personal information .
Approach: They propose to formalise three authentication tasks and their evaluation protocols . they propose to use a spoken multilingual dataset with 5,506 spoken dialogues .
Outcome: The proposed models set the first competitive benchmarks and set directions for future research.
Benchmarking Generation and Evaluation Capabilities of Large Language Models for Instruction Controllable Summarization (2024.findings-naacl)

Copied to clipboard

Challenge: Recent studies have found that large language models (LLMs) can achieve state-of-the-art performance on generic summarization benchmarks, but their performance on more complex summarizing task settings is less studied.
Approach: They benchmark large language models on instruction controllable text summarization . they use 4 evaluation protocols and 11 LLMs to evaluate their performance .
Outcome: The proposed model performs well on instruction controllable text summarization tasks with 4 evaluation protocols and 11 LLMs.
Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search (2025.acl-long)

Copied to clipboard

Challenge: Existing benchmarks and evaluation protocols suffer from inadequate or homogeneous creation of key points, exorbitant cost of data creation, and limited evaluation scopes.
Approach: They propose an automatic framework which leverages Monte Carlo Tree Search to construct numerous and diverse descriptive sentences that thoroughly represent video content in an iterative way.
Outcome: The proposed framework improves MCTS-VCB and DREAM-1K on video captioning tasks by 25.0% and 16.3% respectively.
Reducing Token Redundancy in LVLMs: A Systematic Review of Token Pruning Methods (2026.acl-long)

Copied to clipboard

Challenge: Large Vision-Language Models (LVLMs) excel at visual understanding but face severe computational bottlenecks when processing high-resolution images and long videos due to massive visual token counts.
Approach: They propose a taxonomy categorizing methods into vision-side, LLM-side and hybrid paradigms and analyze token selection mechanisms and pruning strategy.
Outcome: The proposed method selectively removes less informative tokens while maintaining performance.
From Zero to Hero: Cold-Start Anomaly Detection (2024.findings-acl)

Copied to clipboard

Challenge: Existing anomaly detection methods require previous observations to be effective . contaminated observations are often not observed, making them ineffective .
Approach: They propose a method that adapts a zero-shot anomaly detector to contaminated observations . they propose an evaluation suite consisting of evaluation protocols and metrics .
Outcome: The proposed method adapts the zero-shot anomaly detector to contaminated observations.
DrBenchmark: A Large Language Understanding Evaluation Benchmark for French Biomedical Domain (2024.lrec-main)

Copied to clipboard

Challenge: Existing benchmarks for pre-trained language models are limited to only a few languages . a limited number of tasks are evaluated on non-standardized protocols .
Approach: They propose to aggregate diverse downstream tasks into a benchmark to assess PLMs' qualities . they evaluate 8 pre-trained masked language models on general and biomedical-specific data .
Outcome: The proposed benchmark assesses pre-trained language models on 20 diversified tasks.
Beyond Word Boundaries: A Hebrew Coreference Benchmark and an Evaluation Protocol for Morphologically Complex Text (2026.acl-long)

Copied to clipboard

Challenge: CR methods originally designed for English struggle with Morphologically Rich Languages (MRLs) a single token in Hebrew may consist of multiple anaphors, and word/morpheme boundary discrepancies make mention detection and coreference resolution difficult in MRLs.
Approach: They propose a CR dataset that identifies mentions at word, sub-word and multi-word levels and an evaluation protocol that directly addresses word/morpheme boundary discrepancies.
Outcome: The proposed evaluation protocol directly addresses word/morpheme boundary discrepancies in Modern Hebrew, an MRL rich with complex words and pronominal clitics.
Has Machine Translation Achieved Human Parity? A Case for Document-level Evaluation (D18-1)

Copied to clipboard

Challenge: Recent research suggests that neural machine translation achieves parity with professional human translation on the WMT Chinese–English news translation task.
Approach: They empirically test neural machine translation on a Chinese–English news translation task . they show human raters prefer human over machine translation when evaluating documents .
Outcome: The proposed method shows that human translators prefer document-level evaluation over machine translation . the results highlight the need to shift towards document- level evaluation as machine translation improves .
MOA: Multi-Objective Alignment for Role-Playing Agents (2026.acl-long)

Copied to clipboard

Challenge: Prior work on role-playing agents relies on supervised fine-tuning or reinforcement learning with scalarized rewards, but these approaches do not address the coordination of multiple reward dimensions during optimization.
Approach: They propose a reinforcement-learning framework that enables multi-dimensional, fine-grained rubric optimization for general RPAs.
Outcome: Experiments on PersonaGym and RoleMRC show that MOA improves multi-dimensional role-playing performance over supervised and standard RL baselines.
Are Embedded Potatoes Still Vegetables? On the Limitations of WordNet Embeddings for Lexical Semantics (2023.emnlp-main)

Copied to clipboard

Challenge: Knowledge Base Embedding (KBE) models are widely used to encode structured information from knowledge bases, including WordNet, but the evaluation task is often focused on link prediction, ignoring their semantic capabilities.
Approach: They propose to evaluate the performance of Knowledge Base Embedding (KBE) models of WordNet on link prediction and their ability to encode semantic information.
Outcome: The proposed model performs poorly on two semantic tasks and two downstream tasks.
CrossFit: A Few-shot Learning Challenge for Cross-task Generalization in NLP (2021.emnlp-main)

Copied to clipboard

Challenge: We study whether and how cross-task generalization ability can be acquired . we use CrossFit to standardize seen/unseen task partitions and evaluation protocols .
Approach: They propose a problem setup for studying cross-task generalization ability which standardizes seen/unseen task partitions and data access during different learning stages.
Outcome: The proposed model can be used to build few-shot learners across diverse tasks.
How Sampling Affects the Detectability of Machine-written texts: A Comprehensive Study (2025.findings-emnlp)

Copied to clipboard

Challenge: Recent detectors report near-perfect accuracy, often boasting AUROC scores above 99%, but these claims typically assume fixed generation settings, leaving open the question of how robust such systems are to changes in decoding strategies.
Approach: They examine how sampling-based decoding impacts detectability with a focus on how subtle variations in a model’s (sub)word-level distribution affect detection performance.
Outcome: The proposed framework systematically examines how sampling-based decoding impacts detectability, with a focus on how subtle variations in a model’s (sub)word-level distribution affect detection performance.
ReIFE: Re-evaluating Instruction-Following Evaluation (2025.naacl-long)

Copied to clipboard

Challenge: Existing evaluations of large language models (LLMs) for instruction following are incomplete.
Approach: They propose to use 25 base LLMs and 15 recently proposed evaluation protocols to evaluate instruction following on 4 human-annotated datasets.
Outcome: The proposed evaluations identify the best-performing base LLMs and evaluation protocols with a high degree of robustness.
A Survey on LLM-powered Agents for Recommender Systems (2025.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models have demonstrated remarkable capabilities in natural language understanding, reasoning, and generation.
Approach: They present a comprehensive synthesis of large language models and their applications . they dissect a four-module agent architecture and review representative designs .
Outcome: The proposed models address fundamental challenges in traditional recommender systems . they include limited comprehension of complex user intents, insufficient interaction capabilities .
CITB: A Benchmark for Continual Instruction Tuning (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for instruction tuning do not leverage the rich natural language instructions.
Approach: They propose to use a benchmark to study how instruction tuning works in CL tasks.
Outcome: The proposed method can achieve similar or better results than existing CL methods.
On “Human Parity” and “Super Human Performance” in Machine Translation Evaluation (2022.lrec-1)

Copied to clipboard

Challenge: In this paper, we reassess claims of human parity and super human performance in machine translation.
Approach: They reassess claims of human parity and super human performance in machine translation . they argue that human translation involves much more than what is embedded in automatic systems .
Outcome: The proposed results show that human translation involves much more than what is embedded in automatic systems.
Towards Understanding Sample Variance in Visually Grounded Language Generation: Evaluations and Observations (2020.emnlp-main)

Copied to clipboard

Challenge: A major challenge in visually grounded language generation is to build robust benchmark datasets and models that can generalize well in real-world settings.
Approach: They propose to use visual attention to build robust benchmark datasets and models that can generalize well in real-world settings.
Outcome: The proposed models show that human-generated references vary drastically in different datasets/tasks, revealing the nature of each task.
On the Evaluation of Speech Foundation Models for Spoken Language Understanding (2024.findings-acl)

Copied to clipboard

Challenge: Spoken language understanding evaluation (SLUE) benchmarks are used to benchmark complex spoken language understanding tasks on natural speech.
Approach: They propose a set of benchmark tasks to evaluate spoken language understanding on natural speech . they use pre-trained speech foundation models to evaluate the utility of different SFMs .
Outcome: The proposed framework outperforms pre-trained speech foundation models on natural speech . the proposed framework also outperformed self-supervised SFMs on the sequence generation tasks .
LLMs as Narcissistic Evaluators: When Ego Inflates Evaluation Scores (2024.findings-acl)

Copied to clipboard

Challenge: Existing evaluation metrics for natural language generation tasks favor text generated by different LMs . human evaluation by experts is the most reliable approach, but it is costly and time-consuming .
Approach: They examine whether language model-driven evaluation metrics exhibit bias toward underlying language models in the context of summarization tasks.
Outcome: The proposed evaluation metrics tend to assign inflated scores to outputs generated by the very model they are based on.
Structured and Abstractive Reasoning on Multi-modal Relational Knowledge Images (2026.findings-acl)

Copied to clipboard

Challenge: Existing studies on understanding and reasoning with abstractive information from the visual modality have not explored the use of STructured and Abstractive Reasoning (STAR) on such data.
Approach: They propose an automatic STAR data engine to synthesize images with MMRK to build multi-modal instructions with reliable chain-of-thought thinking for various STAR tasks.
Outcome: The proposed framework outperforms GPT-4o in STAR and improves performance across 8 open-source MLLMs.
Your Reasoning Benchmark May Not Test Reasoning: Revealing Perception Bottleneck in Abstract Reasoning Benchmarks (2026.acl-long)

Copied to clipboard

Challenge: Abstraction and Reasoning Corpus and ARC-AGI are widely used to assess progress in artificial intelligence.
Approach: They propose a two-stage pipeline that separates perception and reasoning . they propose to test this pipeline against standard end-to-end one-stage evaluation .
Outcome: The proposed pipeline separates perception and reasoning, and isolates reasoning from bottlenecks.
Defining Knowledge: Bridging Epistemology and Large Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Existing literature on large language models (LLMs) define knowledge as a fact if it correctly completes a cloze sentence . but the predictions of semantically equivalent clozing sentences are inconsistent .
Approach: They review standard definitions of knowledge in epistemology and formalize interpretations applicable to LLMs.
Outcome: The authors compare the preferences of philosophers and computer scientists in terms of knowledge definitions and evaluation protocols for testing knowledge in accordance with the most relevant definitions.
Are Your LLMs Capable of Stable Reasoning? (2025.findings-acl)

Copied to clipboard

Challenge: Existing evaluation protocols and metrics do not capture the full spectrum of LLM capabilities, especially in complex reasoning tasks.
Approach: They propose a new evaluation metric that continuously assesses model performance across multiple sampling attempts, quantifying both the model’s potential capabilities and operational consistency.
Outcome: The proposed evaluation metric measures model performance across multiple sampling attempts and provides comprehensive insights into their potential capabilities and operational consistency.
Stop Playing the Guessing Game! Evaluating Conversational Recommender Systems via Target-free User Simulation (2025.findings-emnlp)

Copied to clipboard

Challenge: despite advances in CRSs, reliably assessing their ability to elicit preferences remains a challenge.
Approach: They propose a user-CRS evaluation protocol with target-free user simulators . they show that current evaluation metrics emphasize single-turn recall of target items .
Outcome: The proposed evaluation protocol is based on a simulation-based evaluation environment.
CLaS-Bench: A Cross-Lingual Alignment and Steering Benchmark (2026.findings-acl)

Copied to clipboard

Challenge: Understanding and controlling behavior of large language models (LLMs) is an important topic in multilingual NLP.
Approach: They propose a lightweight parallel-question benchmark for evaluating language-forcing behavior in large language models across 32 languages.
Outcome: The proposed benchmark measures language steering in 32 languages across 32 languages.
RoboVox: A Single/Multi-channel Far-field Speaker Recognition Benchmark for a Mobile Robot (2024.lrec-main)

Copied to clipboard

Challenge: In this paper, we introduce a new far-field speaker recognition benchmark called RoboVox.
Approach: They introduce a new far-field speaker recognition benchmark called RoboVox which measures the far-feet of a French corpus recorded by a mobile robot.
Outcome: The proposed benchmarks show a significant decline in far-field speaker recognition and urge the community to further research in this domain.
False Sense of Security: Why Probing-based Malicious Input Detection Fails to Generalize (2026.findings-acl)

Copied to clipboard

Challenge: Recent work has leveraged probing-based approaches to study the separability of malicious and benign inputs in Large Language Models’ internal representations.
Approach: They propose to use probing-based methods to study separability of malicious and benign inputs in LLMs' internal representations to detect harmful and benign content.
Outcome: The proposed methods show that they learn superficial patterns rather than semantic harmfulness.
TIU-Bench: A Benchmark for Evaluating Large Multimodal Models on Text-rich Image Understanding (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing text-rich image understanding benchmarks lack scale and fragmented scenarios . a new full-image structured output format is proposed to enable fine-grained evaluation of perception and reasoning capabilities.
Approach: They propose a large-scale, multilingual benchmark that includes over 100,000 annotations and 22,000 question-answer pairs.
Outcome: The proposed framework provides a comprehensive platform for developing and evaluating next-generation multimodal AI systems.
E2EDev: Benchmarking Large Language Models in End-to-End Software Development Task (2026.acl-long)

Copied to clipboard

Challenge: Existing E2ESD benchmarks are limited by coarse-grained requirement specifications and unreliable evaluation protocols.
Approach: They propose a benchmark to assess whether generated software meets user needs . they use a fine-grained set of user requirements and a fully automated testing pipeline .
Outcome: E2EDev is a benchmark to assess whether generated software meets user needs through mimicking real user interactions.
CompassVerifier: A Unified and Robust Verifier for LLMs Evaluation and Outcome Reward (2025.emnlp-main)

Copied to clipboard

Challenge: Existing approaches lack robustness to handle complex edge cases and generalizability across different domains.
Approach: They develop an accurate and lightweight verifier model for evaluation and outcome reward that matches unstructured outputs against standard answers.
Outcome: The proposed model can process multiple answer types including multi-subproblems, formulas, and sequence answers while identifying abnormal/invalid responses.
How Long Reasoning Chains Influence LLMs’ Judgment of Answer Factuality (2026.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) are increasingly adopted as scalable judges for open-ended generation, yet how they form judgments remains insufficiently understood.
Approach: They show that exposing reasoning influences LLM-based judgment . they also show that reasoning fluency and factuality critically shape judgment outcomes .
Outcome: Empirical results show that the presence of reasoning significantly alters judgment behavior . stronger judges exhibit more selective behavior and achieve higher judgment accuracy .
OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Models (2026.acl-long)

Copied to clipboard

Challenge: Existing multimodal reasoning benchmarks for large vision-language models emphasize single-image analysis and fail to exploit contextual information across multiple images.
Approach: They propose a benchmark to evaluate Olympiad-level reasoning when evidence is distributed over multiple images.
Outcome: The proposed model outperforms existing models on bi-image Olympiads and Gemini-3-Pro on multimodal Olympiad-level reasoning tasks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations